Medical Decision Making
○ SAGE Publications
Preprints posted in the last 90 days, ranked by how well they match Medical Decision Making's content profile, based on 12 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.
Epling, J. W.; King, M. J.; Rockwell, M.; Tegge, A. N.; Hester, C. M.; Clay, T. L.; Callen, E. F.; Turner, J. K.; Stein, J.
Show abstract
Introduction: Primary care clinicians (PCC) commonly make decisions in the context of time delay and uncertainty. Delay discounting (DD) and probability discounting (PD) are cognitive biases related to delay and uncertainty that are minimally explored in PCC. We assessed DD and PD in PCC and evaluated their association with low-value care (LVC) decision-making. Methods: We administered a survey to PCC in a Southeastern U.S health system and within the American Academy of Family Physicians networks. The survey comprised standardized psychometric assessments of DD and PD and four LVC clinical vignettes. Outcomes included DD and PD discounting rates for two monetary rewards ($100 and $10,000) and ratings of LVC likelihood (0-100). We used regression analysis with model selection to evaluate the relationship between variables. Results: 225 PCC (89% physicians, 11% advanced practice providers) participated. Heterogeneity in DD and PD rates was observed. For the $10,000 reward, ln k(DD)= -6.80, IQR:-7.60--6.10) and ln h(PD)= 1.75, IQR:1.75-2.36). The reward amount impacted DD and PD in opposing directions (i.e., lower DD/higher PD rates for $10,000 vs. $100). LVC likelihood was highest for low-value antibiotics and lowest for low-value cervical cancer screening (median 20, IQR:10-40 and 0, IQR:0-10, respectively). Model selection revealed demographic associations with LVC likelihood, but no association with DD or PD. Conclusions: Consistent with effects previously reported in non-clinicians, PCC exhibited a range of DD and PD, which ranged by reward magnitude. Neither DD nor PD predicted vignette-based LVC likelihood. Further research should investigate actual clinical practice patterns and other LVC scenarios.
McCready, T.; Thorpe, L.; Roy, B.; Renson, A.
Show abstract
Community-level estimates of healthcare utilization are essential for identifying inequities, allocating resources, and evaluating place-based interventions. However, in the United States, no single data source adequately captures healthcare utilization within geographically defined populations. Population-based surveys often lack sufficient geographic resolution, insurance claims represent only covered populations, and electronic health records are limited to care delivered within participating health systems. Increasingly, researchers combine these fragmented data sources, yet limited guidance exists for conducting valid population-based descriptive analyses using incomplete and overlapping data. We review the strengths and limitations of major data sources used to characterize community healthcare utilization and propose an approach for conducting population-based descriptive analyses using fragmented data. Rather than focusing on the limitations of individual data sources, our approach begins by explicitly defining the target population and the ideal observational study that would answer the research question. Available data sources are then conceptualized as incomplete or imperfect realizations of that ideal, providing a structured approach to (a) identifying sources of selection bias, missingness, and measurement error, (b) articulating required assumptions, and (c) selecting appropriate analytic strategies. We illustrate our approach using colorectal cancer screening utilization among adults residing in Brooklyn, New York during 2022. By shifting attention from individual data sources to the target community and the assumptions required for valid inference, this approach provides a practical approach for strengthening descriptive analyses of community healthcare utilization and informing place-based public health research, policy, and practice.
Lounkaew, K.
Show abstract
National digital health platforms are scaling faster than the evidence on how to finance them. This paper develops a welfare-simulation framework that converts a published willingness-to-pay (WTP) distribution into a prescriptive pricing recommendation, applied to Thailands KhunLook maternal-and-child-health application. Predicted WTP values at the 25th, 50th and 75th unconditional quantiles and the OLS mean -- drawn from a survey of n = 680 Thai parents and relatives of young children previously reported in Lounkaew et al. (2025) -- enter the simulation as parametric inputs. Quintile-level WTP is imputed by monotone-cubic interpolation, a population of 250,000 caregivers is drawn from truncated-Normal distributions around the quintile means, and five financing scenarios are compared: full public provision (S1), a flat market-priced fee (S2), freemium (S3), fine-grained income-tiered pricing (S4), and a means-tested subsidy with a flat fee for the top 60% (S5). A thematic reading of Thai digital-health policy documents bounds the institutionally feasible scenario set and anchors the interpretation of the simulation numbers. Full public provision maximises total welfare at 437.4 million THB but runs a five-year fiscal deficit. The means-tested subsidy gives up about 15% of that welfare to recover 198.6 million THB in net producer surplus, distributes consumer surplus toward lower-income quintiles (concentration index -0.258), and plugs into the existing Thai state welfare card register at near-zero marginal administrative cost. The ranking holds across all twelve sensitivity specifications. Administrative simplicity in subsidy targeting, read against the Thai WTP distribution, dominates finer-grained tiering on both welfare and equity grounds. The framework transfers cleanly to other middle-income countries deciding how to price a national digital health platform. Author summaryMany middle-income-country governments now run free national smartphone apps for the health of mothers and young children, but the funding model is increasingly fragile as initial donor and research grants run out. The question this paper asks is simple: if such a platform had to start charging, what pricing structure would raise the most money without locking out the families who need the app most? Using a published Thai survey of 680 parents and relatives of young children, the paper simulates five alternative designs -- free, flat fee, freemium, fine-tiered by income quintile, and a means-tested subsidy -- and finds that offering the bottom 40% of households free access while charging the top 60% a flat 395 Thai baht per year (roughly USD 11) captures 85% of the welfare of the status-quo free model, generates 199 million baht of fiscal surplus over five years, and distributes benefits toward lower-income users rather than toward the well-off. The design works because Thailands state welfare card register already identifies the low-income target population, so means-testing is essentially free to administer. Other countries with comparable social registries can apply the same logic to their own digital health platforms.
Wang, R.; Chen, H.; Wu, Y.; Li, Z.; Shen, R.; He, F.; Zhao, S.; Zheng, N.
Show abstract
Objective: Chronic care requires sequential treatment under competing biomarker, safety, and cost constraints, yet clinical goal structures differ across diseases. We asked whether one physiology-informed reinforcement learning (RL) paradigm adapts to heterogeneous chronic-care goals without disease-specific policy architectures. Materials and Methods: We formalized a Type A/B/C clinical goal taxonomy (target cure, stable cruise, cycle completion) as a Physiology-Informed Markov Decision Process registry for gout, chronic kidney disease (CKD), and PCOS-mediated fertility treatment--each with PK/PD transitions, discrete actions, safety zones, and guideline doctor baselines. Unified BC->PPO training (GAE lambda=0.95) on 500 simulated trajectories per disease. Evaluation: paired seeds (N=50 primary; N=500 bootstrap 95% CIs), 10-seed robustness, ablation, literature sUA calibration, and out-of-distribution stress. McNemar/Wilcoxon with Benjamini-Hochberg FDR. Results: PCOS (Type C, primary): PPO 72.0% vs. doctor 54.0% at N=50 (+18 percentage points; FDR-significant); at N=500, PPO 69.8% [65.6, 73.8] vs. doctor 52.8% [48.8, 57.2]. Gout (Type A): PPO non-inferior--88.0% vs. 90.0% (McNemar p=1.0). CKD (Type B): doctor 32.0%, BC/PPO 38.0%. Offline CQL 92.0% on gout trajectories. PK recalibration RMSE 97.4 umol/L (r=0.809). Conclusions: Shared BC->PPO training generalizes across three goal types without cross-disease weight sharing. PCOS supports RL for bounded cycles; gout confirms guideline non-inferiority; CKD illustrates cruise-control difficulty. This framework offers a reproducible foundation for chronic pathway optimization pending prospective validation.
Islam, N.; Luo, C.; Tong, J.; Weller, G.; Polleya, D. A.; Kent, A.; Bair, S.
Show abstract
Introduction In analyses of time-to-event data, clinical characteristics can have non-linear impacts on survival outcomes, and understanding this dynamic behavior is crucial for producing real-world evidence (RWE). Nonetheless, estimating these dynamic effects is inherently challenging when utilizing real-world data (RWD), especially since sharing individual-level patient data (IPD) is heavily restricted due to regulatory limitations. Additionally, computational difficulties are exacerbated by the high dimensionality, inter-dependency, rarity, sparsity, and scarcity of features. While data augmentation through collaboration across multiple sites might address these challenges, such collaboration is often infeasible and hindered by regulatory measures that protect patient privacy, thereby preventing the sharing of IPD between sites. Objectives To address this challenge, we propose a privacy-preserving regularized algorithm that eliminates the necessity of aggregating any protected health information across sites. This algorithm employs a penalized federated additive model utilizing piecewise exponential survival (FAMES) data and estimates non-linear effects of features while accounting for non-varying confounding effects. The model is flexible and can accommodate both multiple and multivariate smooth effects simultaneously. Methods The proposed model transforms survival data into a piecewise exponential data (PED) structure and casts the semi-parametric optimization problem into a generalized additive modeling framework assuming Poisson distribution. The model uses orthonormal splines to approximate non-linear effects and incorporates L2-norm based penalty terms to control the smoothness and goodness-of-fit of these effects. The algorithm is optimized using site-specific aggregated summary statistics and is solved iteratively through the Newton-Raphson method. Results The model is employed to assess the smooth effects of clinical features, such as age and numeric laboratory values, on overall survival using RWD from approximately 874 newly diagnosed Acute Myeloid Leukemia (AML) patients treated at seven distinct sites in the United States. The model exhibited non-linear smooth effects for lactate dehydrogenase, platelets, and others underscoring their strong association with disease prognosis. The model demonstrates a lossless property, providing estimates of smooth and fixed effects that are comparable to those derived from the pooled PED. Additionally, the inference of parameters for testing the nullity of effects remains consistent. This model is communication-efficient, necessitating roughly twelve rounds of communication across sites. Conclusion We anticipate that this model can facilitate multisite collaboration and enable smaller sites to participate in generating and validating RWE, especially for rare diseases. While the model was applied within the context of AML, it is disease-agnostic and can be implemented in any other clinical context and across various sites globally without losing any generality.
Ha, Y.; Park, H.; Lee, Y.; Kim, S.; Ahn, S.
Show abstract
BackgroundDisability weights (DWs) quantify the severity of health loss and are essential for estimating disability-adjusted life years in the Global Burden of Disease (GBD) framework. Conventional DW estimation relies on resource-intensive population surveys that are difficult to update or adapt to emerging health states. Large language models (LLMs) may offer a scalable alternative by approximating human perceptions of disease severity through structured judgment tasks. MethodsThis exploratory study evaluated the alignment between LLM-derived and human-derived DW rankings using 222 health states from GBD 2010. All possible pairwise comparisons (24,531 pairs, each repeated three times) were conducted across four LLMs (GPT-5 mini, GPT-5, Claude Haiku 4.5, and Claude Sonnet 4.5). DWs were estimated via probit regression and evaluated using Spearmans rank correlation and Steigers z test. The effects of prompt language (English vs. Korean), cultural role prompting, and medical specialist role prompting on alignment were examined. Additionally, the Binomial-Logit Indifference-Point (BLIP) estimator was proposed and validated through leave-one-out cross-validation for estimating DWs for health states without established values. ResultsAll four LLMs showed high rank correlation with GBD 2010 DWs (Spearmans {rho} = 0.893 to 0.909), with no significant inter-model differences. Korean-language prompting significantly improved alignment with Korean DWs ({rho} = 0.756 vs. 0.715, p = 0.011), and Korean cultural role prompting improved alignment with both GBD 2010 DWs ({rho} = 0.922 vs. 0.909, p = 0.002) and Korean DWs ({rho} = 0.738 vs. 0.715, p = 0.001). Medical specialist role prompting significantly reduced alignment with GBD 2010 DWs ({rho} = 0.895 vs. 0.909, p = 0.001). BLIP demonstrated strong agreement with GBD 2010 DWs (Pearsons r = 0.862, MAE = 0.066) and produced plausible estimates for Long COVID (mild: 0.020, moderate: 0.298, severe: 0.529). ConclusionsLLMs can approximate human perceptions of disease severity with high rank-order consistency. Prompt language and role framing significantly influenced alignment, with culturally grounded lay prompting enhancing and specialist prompting reducing correspondence with population-based DWs. BLIP provides a practical framework for generating provisional DW estimates for emerging or underrepresented health states when conventional surveys are infeasible.
Adibi, A.; Le, K. X.; Pierson, E.; Diao, J. A.; Esfandiari, N.; Carlsten, C.; Sadatsafavi, M.
Show abstract
Importance: Several professional medical societies have removed race and ethnicity from widely used clinical algorithms with implications for millions of patients. Yet the opinions of patients and the public regarding the tensions underlying these pivotal changes have not been systematically explored. Objective: To assess global public opinion on the use of race or ethnicity in clinical algorithms, including preferences for different approaches to algorithmic reform and perceptions of alternative predictors. Design: Cross-sectional survey study. Setting: Multinational opt-in online survey conducted via Prolific in January 2026. Participants: A volunteer convenience sample with quota sampling to achieve approximately equal participation by sex at birth and across ten categories of self-identified race and ethnicity. Main Outcomes and Measures: Self-reported comfort with demographic and social predictors in clinical calculators, with net comfort defined as percentage extremely or somewhat comfortable minus percentage extremely or somewhat uncomfortable; preferences for race-specific versus race-free algorithms; perceptions of algorithmic harm or benefit. Results: Of 1,050 responses, 994 (94.7%) met eligibility criteria. Participants resided in 43 countries with a median age of 32.0 years (IQR, 26-41). Net comfort with the use of race or ethnicity in a hypothetical cancer risk calculator was +62.4% (95% CI: +57.8% to +66.9%), compared with +14.5% (95% CI: +9.1% to +19.9%) for postal or ZIP code. Overall, 87.9% (95% CI: 85.9% to 90.0%) were comfortable with race or ethnicity if a clinician explained its use and only 12.8% agreed race and ethnicity should never be used clinically. Across spirometry, kidney function, and cardiovascular risk calculators, 40.0% to 47.6% preferred race-specific versions, whereas 16.7% to 28.2% preferred race-free alternatives. Furthermore, a substantial proportion disagreed that they were well-represented by race and ethnicity categories, ranging from 22.1% for osteoporotic fracture risk equations to 42.9% for cardiovascular risk equations. These findings were consistent across countries, self-identified race and ethnicity, and among participants reporting prior experiences of racism in healthcare. Conclusions and Relevance: In our diverse multinational survey study, respondents were comfortable with the use of race and ethnicity across application areas, but often did not feel represented by existing categories and were less comfortable with the use of alternatives based on postal or ZIP codes.
Kim, D.; Pasco, R.; Johnson, K. E.; Fox, S. J.; Reich, N. G.; Meyers, L. A.
Show abstract
Accurate outbreak forecasts are critical for timely and effective public health response. In the United States, however, most forecasts are produced at the state level, which can mask substantial sub-state heterogeneity and limit their utility for local planning. We generated and evaluated forecasts of the percentage of Emergency Department visits attributable to influenza across 173 large metropolitan Health Service Areas (HSAs) using a gradient boosting quantile regression (GBQR) model, and compared their accuracy to forecasts derived from state-level data alone. At a one-week, two-week and three-week horizon, local forecasts outperformed state-based forecasts in 98.8%, 90.8%, and 78.6% of HSAs, respectively, achieving mean weighted interval scores that were on average a 39.2% lower (95% range: 5.9% to 76.7%), 19.6% lower (-6.3% to 59.5%) , and 11.4% lower (-11.7% to 44.9%), respectively. The performance advantage of local forecasting was strongest in HSAs representing a smaller share of their state's population and increased with the proportion of the HSA population living in urban areas and the number of metropolitan areas within a state. These results, based on an analysis of HSAs with populations greater than 250,000, demonstrate that fine-scale modeling can substantially improve forecast accuracy and highlight the potential value of local forecasts for outbreak preparedness and response.
Chizari, H.; Peter, N.; Lin, B.; Malekinezhad, F.; Pietroni, M.
Show abstract
Elective surgery late cancellations and ``did not attend'' (LCDNA) events waste theatre capacity, lengthen waiting lists, and impose avoidable costs on NHS Trusts. We present a decision-support approach that ranks upcoming elective procedures by expected cancellation cost and supports capacity-constrained outreach by selecting the highest-risk Top-K cases for intervention. Using cost-sensitive learning and a clinically grounded cost model, the policy reduces expected cost from approximately 103 GBP per case under business-as-usual to 77.08 GBP per case in a hospital-holdout (cross-site) evaluation designed to mimic deployment to a new hospital. In a complementary time-forward evaluation, representing prospective use within the same service environment, expected cost falls further to 70.97 GBP per case. The 6.11 GBP per-case difference between the two regimes highlights the added uncertainty introduced by cross-site operational shift and supports a conservative roll-out with local calibration and monitoring. Explainability analyses suggest that booking-to-procedure lead time, specialty or service line, calendar effects, and prior cancellation history are the strongest drivers of prediction, helping to inform tiered intervention workflows that prioritise near-term bookings and use model--pathway mismatches as an audit signal. Overall, the framework turns predictive performance into practical, capacity-aware policy guidance for reducing avoidable cancellations while supporting safe and equitable implementation.
Fitch, K. V.; Santaularia Gomez, N. J.; Tanveer, M.; Holmes, G. M.; Moracco, K. E.; Fliss, M. D.; Fulcher, N.; Ranapurwala, S. I.
Show abstract
Introduction: Even though state minimum wage (MW) is a policy lever that affects income and poverty and can prevent of violence, no prior study has comprehensively evaluated its impact in the United States (US). In this study, we estimated the impact of at least a $1 USD increase in state MW above the federal MW on overall, firearm, and non-firearm homicide mortality and examined its impact on racialized inequities. Methods: We conducted a quasi-experimental study using controlled interrupted time series (CITS) and synthetic controlled interrupted time series (SCITS) approaches to examine immediate and sustained impact of state MW increases. We used state-month level homicide victimization mortality data from 2010-2019. Homicide victimization death was identified using International Classification of Disease codes, 10th revision. State MW data was obtained from the Bureau of Labor Statistics. Results: Demographic and social variables from intervention, never-exposed, and always-exposed states were similar to each other and representative of the total US population from all 50 states. The CITS results show that after MW increases in the intervention states, these states experienced a sustained decline of -0.22 (-0.37, -0.07) homicide victimizations/ 100,000 person-years/ year relative to the never-exposed states and -0.39 (-0.59, -0.18) relative to always-exposed states. This resulted in 5,657 fewer homicide victimization deaths in the intervention states over four years of post-MW increase period compared to the never-exposed states. SCITS results were similar to the CITS results, and the majority of sustained declines were observed in firearm-related deaths and among Black population. Conclusion: MW increase was associated with a reduction in homicide victimization rates, which were robust in multiple sensitivity analyses, more pronounced for firearm-related homicide deaths, and reduced homicide victimization inequities for Black Americans.
Lee, A.; Kazemi, S.; Wilson, P.; Thaker, K.; Kwan, L.; Cabri, J.; Li, K.; Dunn, M.; Yaghoubian, A.; Elkhoury, F.; Scotland, K.; Saigal, C.
Show abstract
Introduction Patients with nephrolithiasis face challenges in making a high-quality, preference sensitive decision. Our prior work established feasibility and patient acceptance of a software-based decision aid (DA). The objectives for this study were to identify implementation strategies for the DA in routine care and determine whether DA implementation enhances decisional quality for patients. Methods New nephrolithiasis patients were recruited from the institution Medical Center from June 2018 to April 2024 to receive a software-based pre-visit DA that measured care preferences and used decision analysis to rank treatments. The RE-AIM framework and Plan-Do-Study-Act (PDSA) cycles were used to improve implementation outcomes. Patients completed survey instruments evaluating decisional conflict, shared decision-making, care satisfaction, and treatment choice following their provider visit. These metrics were compared in the DA cohort (n=81) to those in a usual care cohort (n=78) with Wilcoxon rank-sum and Chi-square (or Fishers exact) tests. Results Implementation data revealed sustained reach and progressive improvement in fidelity. The DA cohort reported higher decisional quality relative to controls (p=0.003) and reported greater support/advice to make a choice (p=0.005). The DA cohort more often discussed options with their doctor (87.5% vs 69.2%, p=0.005) and were more likely to be promoters of their provider (p<0.001) and health system (p=0.029). The DA cohort was less likely to have switched their treatment preference post-consultation (32.1% vs 71.8%, p<0.001) suggesting greater consistency in decision-making. Conclusions Software-based DAs in nephrolithiasis can mitigate decisional conflict, improve SDM, and improve patient satisfaction. Further work should explore broader implementation and long-term clinical outcomes.
Velasco Pardo, V.; Daines, L.; Katikireddi, S. V.; Ritchie, L.; Robertson, C.; Simpson, C. R.; McCowan, C.; Swallow, B.
Show abstract
Background During the COVID-19 pandemic, public health agencies used near real-time observational data to answer questions regarding vaccine effectiveness. However, traditional observational methods do not allow conclusions regarding counterfactual scenarios to be drawn from clinical data. Counterfactuals, which are outcomes that would have occurred under alternative interventions, can be used to formally assess the causal effects of public health interventions on health outcomes while accounting for the effects of confounding. Ideally individual patient data is used for the development of counterfactuals. Low-fidelity synthetic data may be useful for advancing methodological development where governance and privacy constraints prohibit access to sensitive personal data. Methods We simulated synthetic datasets based on the EAVE-II COVID-19 platform which has been limited to use for surveillance purposes. EAVE-II includes almost all resident people in Scotland registered with qualified general medical practitioners. Patient characteristics were simulated to reflect the known distribution of the Scottish population, accounting for dependencies between variables. Each synthetic dataset was encoded to different realistic scenarios for EAVEII 'ground truth' vaccine rollout and effectiveness results, explicitly stating the causal and confounding mechanisms, using a statistically sound method based on marginal structural models. Synthetic datasets of 100,000 individuals were then generated across five confounding scenarios and five severe outcome types. Results In scenarios with weak confounding, both unweighted and inverse probability of treatment weighted (IPTW) logistic regression recovered the true causal parameters. As confounding strength increased, only weighted models recovered the true mechanism. Conclusions Low-fidelity synthetic datasets simulated from EAVE-II data analysts to build and test causal inference pipelines, develop novel analysis pipelines, and train new researchers while awaiting access to real data. We showed how to generate synthetic datasets from a marginal structural model under different confounding scenarios.
Conde, F.
Show abstract
Background: Health-related social needs (HRSNs), particularly housing instability, are significant drivers of poor health outcomes among Medicaid populations. New York State's Social Care Networks (SCNs) aim to systematically connect members to housing services through coordinated referral systems. However, limited systematic analysis of referral patterns hinders quality improvement efforts. We analyzed housing referral outcomes and workflows to identify barriers to successful service connections. Methods: We conducted a mixed-methods quality improvement study at Public Health Solutions' WholeYouNYC SCN Coordination Center. Quantitative analysis examined 4,258 housing referrals submitted between June 2025 and January 2026, extracted from the Unite Us platform via Power BI dashboard. We calculated acceptance rates, analyzed time metrics, and examined outcomes by receiving organization. Qualitative data were collected through structured consultations with 7 staff members (5 navigators, 2 supervisors) and review of internal workflow documentation. Process mapping identified workflow bottlenecks. Results: Of 4,258 housing referrals, only 45% (n=1,936) were accepted by receiving organizations, while 19% (n=815) were rejected and 32% (n=1,382) remained awaiting response with no recorded action. Average time to acceptance was 8 days for accepted referrals. Acceptance rates were consistent across top receiving organizations (44-46%), suggesting systemic rather than partner-specific barriers. Analysis of unresolved referrals revealed prolonged cases, with the longest pending 271 days. Three critical workflow bottlenecks were identified: CBO response delays, missing housing documentation, and challenges with client engagement. Conclusions: Low housing connection rates (45%) and prolonged unresolved referrals (up to 271 days) indicate systemic barriers requiring interventions at multiple levels. Recommendations include establishing CBO response time benchmarks, implementing automated follow-up protocols, standardizing documentation requirements, and enhancing real-time data monitoring. These findings provide an evidence-based framework for quality improvement in social care coordination programs.
Henry, K.; Blotske, K.; Smith, B.; Li, T.; Gao, Y.; Zhao, X.; Liu, T.; Sikora, A.
Show abstract
Background: Standardized evaluation of agentic artificial intelligence (AI) for medication management is lacking. Given the potential lethality of medication errors endorsed or missed by AI, performance evaluation constructs are essential. The purpose of this evaluation was to develop a standardized grading framework for performance evaluation of medication management tasks. Methods: A mixed-methods approach was undertaken that included literature evaluation for standards and best practices of comprehensive medication management (CMM), panel discussions, and iterative application to set of cases. The goal was to develop a grading framework that effectively evaluated domains like safety, factuality, and clinical relevance that can be employed for a broad range of medication domains (i.e., electrolyte replacement, antibiotic selection). Inter-rater reliability with intraclass Krippendorffs Alpha was the primary outcome. Results: A total of 5 panelists developed the CMM Evaluation Framework, which includes 4 dimensions: safety, factuality, completeness, and preference. These dimensions are applied to three CMM skills: collecting patient data, analyzing information, and designing regimens. Each dimension is rated from 1-5. An additional dimension evaluated the presence of hallucinations and errors with high harm scores (i.e., absolute failure criteria regardless of an overall score). The Krippendorffs Alpha was highest in the medication therapy problem and medication therapy format categories, for 50 pneumonia cases, run in triplicate (150 total). Conclusions: This framework is informed by national standards for CMM and the healthcare professionals dedicated to the provision of this service. These domains allow for the possibilities of practice variation via the preference domain while also having strong guardrails against the commission of medication errors. Further analyses beyond pilot testing are necessary.
Mandke, C.; Agrawal, H. K.; Bharti, B.; Chansoria, M.; Gupta, G.; Rawat, S. K.; Sarkar, N. K.; Singh, A.; PS, S.; Walia, S.; VALID (Validation of AI in Low-resource and Indian Domains) Consortium,
Show abstract
BackgroundHealthcare providers in low- and middle-income countries (LMICs) are increasingly relying on Artificial Intelligence (AI) tools, yet most available AI assistants are general-purpose systems not designed for the specific clinical, epidemiological, and resource contexts of these settings. There is no evidence, from physicians assessments, on whether clinical reasoning support from purpose-built, context-specific and retrieval-augmented AI tools can outperform general-purpose AI agents. MethodsWe conducted a prospective multi-site validation study enrolling 37 physicians across India and Bangladesh. Each physician evaluated two AI tools (a) VITA (Validated Intelligence for Treatment and Assessment), a purpose-built (context-specific and retrieval-augmented) clinical reasoning AI assistant trained on India-specific guidelines, antimicrobial resistance patterns, and formulary constraints, and (b) ChatGPT Plus (version 5.2), a leading general-purpose AI assistant on six hypothetical clinical case vignettes (three predefined, three physician-selected). Evaluations were scored across six dimensions (differential diagnosis, clinical workup, treatment recommendation, dosing, clinical decision-making, and evidence quality) on a 1-5 Likert scale, yielding 444 observations. Analyses included paired t-tests, Wilcoxon signed-rank tests, and multivariate regressions with robust standard errors. ResultsVITA scored significantly higher than ChatGPT across all six evaluation dimensions. The mean composite score (sum of all dimensions, maximum = 30) was 25.4 for VITA versus 22.3 for ChatGPT (difference = +3.1 points, t = 8.31, p < 0.001). The largest advantage was in evidence quality (VITA: 4.46 vs. ChatGPT: 3.14, a 42% relative gap). VITAs advantage was consistent across both predefined and doctor-defined hypothetical cases and was robust to controls for physician demographics, case type, and evaluation order in multivariate regression (coefficient = +3.08, p < 0.001). ConclusionsIn this first systematic head-to-head physician evaluation of a purpose-built clinical reasoning AI assistant versus general-purpose AI in an LMIC setting, physicians consistently rated the context-specific tool as superior. These findings suggest that contextual relevance--including local guidelines, formulary constraints, and resistance patterns--matters for clinical AI adoption and quality in resource-limited settings.
Gensheimer, M. F.; Adhikari, R.; Parmer-Chow, C.; Liu, N.; Ma, S.; Shieh, L.
Show abstract
Background: Manual review of 30-day hospital readmissions can identify actionable quality and safety problems, but it is labor-intensive. We developed and evaluated an agentic AI workflow for evidence-grounded readmission review. Materials and methods: We studied adult patients with unplanned 30-day readmission after discharge from a medicine hospitalist service at a single academic health system. An AI agent using a large language model queried a database containing notes, encounters, procedures, laboratory results, and other clinical data, and completed the same structured readmission-review rubric used by physicians. In the primary comparative evaluation, 20 randomly selected readmissions from 2025 were each reviewed by two physicians and the AI system. Blinded physician evaluators rated review quality. After rubric refinement, the AI workflow was applied to 100 recent readmissions in an exploratory expanded-cohort analysis of recurring improvement opportunities. Results: In the primary comparative evaluation, the AI classified 9/20 readmissions (45%) as preventable, compared with 19/40 physician reviews (47.5%). Blinded overall quality ratings were similar for AI and physician reviews (4.35 vs. 4.20 on a 1-5 scale; mean difference 0.15, 95% CI -0.20 to 0.48; p=0.49), as were factuality/support and usefulness/actionability ratings. No AI hallucinations were identified during factuality review. Agreement on preventability and primary readmission category was low for both AI-human and human-human comparisons. The AI system cost $0.23 per chart; physician reviewers took a median of 15 minutes, corresponding to an estimated $42.43 per chart. In the exploratory expanded-cohort analysis, AI-assisted review identified recurring vulnerabilities in post-discharge follow-up plans, incomplete inpatient workups, medication-safety transitions, and indwelling-device transitions. Conclusions: Agentic AI produced readmission reviews with similar blinded quality ratings to physician reviews in this small single-center primary comparative evaluation and supported identification of recurring quality-improvement themes in the exploratory expanded-cohort analysis. Preventability judgments remained variable among both AI and physicians, underscoring the need for human oversight and prospective evaluation before operational use.
Dojcsak, L.; Abegaz, T.; Islam, M.; Chandler, Y.; Maleku, A.; Doubeni, A.; Mohammed, B.; Langston, M. A.; Donneyong, M. M.
Show abstract
Health-related social needs (HRSNs), such as housing instability, food insecurity, and transportation challenges, are nonmedical factors associated with poorer health and well-being. Screening for unmet HRSNs is a critical step towards identifying at-risk patients, but manual screening is resource intensive and often incomplete. We utilized Electronic Health Records (EHR) data to develop machine learning models to identify unmet HRSNs using a limited set of non-modifiable sociodemographic features available in EHRs. We included 745,975 patients screened for at least one HRSN using data from community health centers that participated in the OCHIN practice-based research network between 2016 and 2022. Logistic regression, random forest (RF), eXtreme Gradient Boosting (XGBoost), and Light Gradient Boosting Machine (LightGBM) algorithms were trained to predict unmet HRSNs. Model performance was evaluated using 10-fold cross-validation and area under the receiver operating characteristic curve (AUROC). For overall HRSN prediction, LightGBM (AUROC, 64.5%, 95%CI: 64.3, 64.7) performed slightly better than logistic regression (61.4%), RF (63.7%), and XGBoost (60.3%). Similar performances were observed predicting individual HRSNs. Model performances were modest; however, they establish a benchmark for predictive performance achievable using only routinely available demographic data and provide a foundation for incorporating additional clinical and area-level social determinants of health data.
Osborne, T.; Mahmud, T.; Zheng, X.; Jampala, S.; Abbasi, S.; Hong, S.; Kranz, K.; Lee, S.; Ng, P.; Odekon, K.; Schachter, L.; Sexton, R.; Spinnato, T.; Tharakan, M.; Wu, Z.; Wang, F.; Wong, R.
Show abstract
Although large language models (LLMs) have shown promise for discharge summary generation, their value may be greater in longer hospitalizations, where increasing documentation volume and complexity increase both clinician burden and the risk of communication failures during transitions of care. Prior evaluations of LLM-generated discharge summaries have largely involved shorter stays and have rarely examined receiving-clinician priorities or incidental finding reporting. We compared LLM-generated and human-authored discharge summaries for 60 Internal Medicine hospitalizations lasting 7 to 21 days, with paired assessment by hospitalists and primary care physicians (PCPs). Clinician reviewers preferred LLM-generated summaries for 95% of encounters and rated them higher for quality, readability, factuality and completeness. PCPs, the primary recipients responsible for post-discharge care, found that LLM-generated summaries were better for understanding and communicating hospital care to patients, and providing follow-up care. LLM-generated summaries had fewer annotated errors, primarily due to fewer omissions, without increased estimated harm potential or likelihood compared with human-authored summaries. Benefits of LLM-generated summaries were especially salient for PCPs, who identified more omissions with greater downstream likelihood of harm than hospitalists. This underscores the importance of designing transition documents around the needs of clinicians assuming care post-discharge. LLM identification of radiology incidental findings was generally accurate and appropriate, suggesting potential to improve follow-up of clinically relevant findings. These findings extend prior work by demonstrating clinical value of LLMs in summarizing longer, complex hospitalizations and highlighting the value of stakeholder-centered design in clinical AI systems. Together, they support supervised LLM-assisted discharge summarization as a tool to reduce cognitive burden, improve documentation quality, and enhance transition-of-care communication.
Di Carluccio, E.; Koliopanos, G.; Ojeda, F. M.; Weimar, C.; Ziegler, A.
Show abstract
Statistical prediction models for binary outcomes are becoming increasingly popular. One significant challenge is calibrating these models to suit the characteristics of a target population that is structurally different from the original population. Calibration is especially challenging when there is no training data available from the target population. To address this problem, we propose a novel calibration method, SimCal, which uses synthetic data generated from the model development data in conjunction with marginal statistics from the calibration cohort. We show that expert judgment modeling (EJM) may be used for calibration if cross-sectional data from the target population are available comprising expert judgments about the potential outcome and the covariates. We describe three alternative calibration approaches when calibration data are lacking: similarity-binning averaging (SBA), adaptive calibration of predictions (ACP), and Elkan calibration. In a simulation study, we compare SBA, ACP, Elkan calibration, and SimCal. R code for applying these methods is provided from the re-analysis of data on coronary artery disease. We illustrate all 5 calibration approaches with a real data set for predicting functional outcome after stroke and all approaches but EJM in the re-analysis of the Cleveland Clinic data. None of the approaches performed convincingly well in all situations. SimCal performed well when model parameters were correctly specified. EJM failed on the stroke data. Further research is urgently required for calibration in the absence of calibration data.
Corona-Moreno, R.; Acuna-Zegarra, M. A.; Santana-Cibrian, M.; Velasco-Hernandez, J. X.
Show abstract
During the COVID-19 pandemic, limited testing capacity and reporting delays complicated epidemic surveillance and decision-making in Mexico. We calibrated \textit{covidestim}, a Bayesian nowcasting model, to estimate the total SARS-CoV-2 infections from reported cases and deaths using Mexican surveillance data. Disease-progression distribution priors were calibrated using Mexico City records and validated through comparisons with national seroprevalence surveys, hospitalization data, and annual reported severe-case rates across all states. Using the reconstructed estimates of active infections, we implemented an event-based risk framework that quantifies the probability of encountering at least one infectious individual in gatherings of different sizes. This probability was subsequently translated into a four-level epidemiological traffic-light indicator and computed at both state and municipality levels. The resulting estimates revealed substantial spatial heterogeneity that is obscured by state-level aggregation, particularly in states with marked differences between urban and rural municipalities. To evaluate consistency with public-health indicators, we compared the proposed risk classification with the official Mexican epidemiological traffic-light system, considering interpretable gathering sizes relevant to public-health decision making. Weekly reports derived from this framework were delivered to policymakers in the State of Queretaro in Mexico, as an anticipation tool for school reopening and public-space management. This demonstrates that this Bayesian reconstruction of infections combined with event-based risk metrics can provide an interpretable and generalizable municipality-level complement to routine surveillance systems, particularly in regions with limited testing capacity and heterogeneous local transmission dynamics.